Skip to content

fix(aiguard): scope framework collision avoidance to request and response phases - #20461

Open
avara1986 wants to merge 20 commits into
mainfrom
avara/aiguard-phase-scoped-collision-avoidance
Open

avara1986 wants to merge 20 commits into
mainfrom
avara/aiguard-phase-scoped-collision-avoidance

Conversation

@avara1986

@avara1986 avara1986 commented Sep 21, 2026 •

Copy link
Copy Markdown
Member

Description

Jira: APPSEC-70286, APPSEC-70474, APPSEC-70480 and APPSEC-70282

AI Guard used a single all-or-nothing depth counter to stop a framework integration and a provider integration from evaluating the same call twice. While it was raised, provider listeners skipped both their request and their response evaluation. LangChain streaming raised it on iteration entry but never evaluated the streamed response, so with DD_AI_GUARD_ANALYZE_STREAM_RESPONSES_ENABLED=true a streamed LangChain response was scanned by nobody.

Phase-scoped claims

The counter is replaced by independent REQUEST and RESPONSE claims, so a framework only suppresses the provider checks it performs itself:

Caller Claims Evaluates
LangChain generate / agenerate REQUEST + RESPONSE, for the whole call request in .before, response in .after (#20457)
LangChain model _stream / _astream REQUEST + RESPONSE, only while the model stream is read (with stream analysis off, only its first read) request in the model-stream buffer before the model is read (skipped inside that model's own generate, whose .before already ran), response in the same buffer (below)
Strands REQUEST + RESPONSE before- and after-model-call hooks

Provider .before listeners check REQUEST; provider .after listeners and the buffered streams check RESPONSE.

AI Guard owns its LangChain logic; the contrib stays product-agnostic

The LangChain contrib (ddtrace/contrib/internal/langchain/) carries no AI Guard state or naming. It only dispatches the product-agnostic langchain.{chatmodel,llm}.{generate,agenerate,stream}.before and .{generate,agenerate}.after events (AI Guard no longer listens to .stream.before; see below), from inside the LLM span, so a listener's block is still recorded on that span and the LLMObs output is kept. Against main the contrib diff is removals only, plus renaming the stream option aiguard_before_event to before_event: the AI Guard state dicts, the aiguard_*_event options, the stream handler's start_stream override and the .finally / .stream.started / .stream.finally events are gone. They existed only so AI Guard could release a claim in a different listener from the one that took it.

Instead, at langchain.patch AI Guard installs its own wrappers (outermost, around the contrib's) on BaseChatModel / BaseLLM:

  • generate / agenerate: with aiguard_context(REQUEST, RESPONSE) around the call.
  • _stream / _astream of each model class: the stream buffer below, which also takes the stream's claim. Installed at langchain.patch for every existing subclass and from an __init_subclass__ hook for classes defined later, so a path that reads _stream directly (stream_events(version="v3")) is covered on first use.

Every claim is taken and released in one frame, so no handle travels through the contrib and nothing can leave a claim behind: not a nested inner stream, not an early break, not a close from another task.

The claim covers only the read of the model's stream, so the caller's loop body and LangChain callbacks are never inside it. This fixes APPSEC-70480: direct OpenAI / Anthropic calls made while iterating a LangChain stream are evaluated. Previously the claim spanned the whole iteration and those calls skipped AI Guard.

Streamed LangChain responses are buffered around each model's _stream / _astream

With DD_AI_GUARD_ANALYZE_STREAM_RESPONSES_ENABLED=true, AI Guard wraps the _stream / _astream each model class defines. The whole model stream is read, the chunks are merged into the message LangChain would build, request + response are evaluated, and only then are the chunks replayed. Subclasses define these methods themselves, so the wrappers are installed per class (see above). The layer matters:

  • Below LangChain's callbacks. stream(), astream_events(), LangGraph stream_mode="messages" and generate's internal streaming all hand each token to callbacks after reading _stream. No callback sees a token before the verdict, and astream_events() still emits its stream events because the model run is still open when the first chunk is replayed. Providers that report tokens to the run manager from inside _stream (LLMs, ChatOpenAI(streaming=True)) get a deferring run manager whose tokens are replayed after the verdict.
  • Streaming inside invoke() (APPSEC-70474). When LangChain streams inside generate (a streaming callback handler, streaming=True, stream=True), the buffer evaluates the response before any token is reported. It records the payload it evaluated, and .generate.after skips that exact payload, so the response is evaluated once. The record is scoped to the generate call, so a cached response is still evaluated.
  • Timeouts. langchain-openai wraps each provider read inside _astream in its own 120 s wait_for. The buffer drains _astream from outside, so every provider read keeps its own timeout.
  • Blocks cannot be swallowed. The buffer raises the BaseException-based AIGuardAbortError. The provider's OpenAIAIGuardAbortError is an Exception, which with_fallbacks and retry policies caught and routed to an unevaluated fallback.
  • No double tool-call evaluation. On a clean verdict the evaluated tool calls are recorded in the last chunk's response_metadata; chunk addition carries it into the aggregate, so the legacy AgentExecutor (which streams by default) and the langgraph hooks skip them.
  • The buffer evaluates the request too. Before claiming and reading the model, it evaluates the request unless that model's own generate call is running (tracked per model, so a different model streamed inside a generate still gets its check). Without this, a direct _stream read (stream_events v3) had no request check while its claim made the provider skip its own, so with the flag off its request was evaluated by nobody. This makes the langchain.*.stream.before listeners redundant; they are removed, so stream(), astream() and direct reads take one path. A blocked stream() now raises on its first next() rather than at the stream() call, and LangChain's on_llm_error callback fires for the block.
  • Native protocol events. langchain-core 1.4 chat models can implement _stream_chat_model_events / _astream_chat_model_events, which stream_events / astream_events v3 read instead of _stream. Those hooks get the same per-class buffer: the request is evaluated first, the events are drained and rebuilt into a message with LangChain's ChatModelStream, evaluated, then replayed, so on_stream_event callbacks see nothing before the verdict. Tool calls on this path are not marked as evaluated, so an agent hook may evaluate them again (one extra evaluation, not a missed one).
  • The provider iterator is always closed. Sync and async, buffered or passthrough, the model's iterator is closed after a full drain, a failure or a cancellation, so class-based SDK streams are not left open.
  • Patched classes stay collectable. The per-class record of wrapped methods is a WeakKeyDictionary, so model classes defined at runtime can be garbage-collected.
  • Any provider. ChatAnthropic and every other chat model or LLM are covered, not only OpenAI. The buffer is the evaluator, so it does not step aside for an outer framework's claim: a LangChain stream inside a Strands tool is evaluated.
  • Cost. One claim per stream instead of one per chunk, and the chunks are merged in one add_ai_message_chunks call instead of pairwise.

It is passthrough when the flag is off. Conversion errors fail open; only a block raises. A subclass _stream that calls super()._stream is buffered once: the buffer skips only a nested read of the model it is already reading, on every read of that model (so delegating after a first chunk is not buffered again), while the claim still covers only the first read. Another model streamed during that read (a router model's inner model) gets its own buffer.

Claims are shared objects, released by handle (APPSEC-70282)

An asyncio task works on a copy of its parent's Context, so a counter lowered from another task left the claiming task covered for good, and later OpenAI/Anthropic calls in it skipped AI Guard. Claims are now shared objects stored in a ContextVar: a release from any task is seen by all, never raises (the old ContextVar.reset raised ValueError into the framework's cleanup path), and never drops an unrelated claim.

Every claim is released by its handle, never by searching the current context for "the latest claim", which misses claims made in another task and takes an inner call's claim when calls nest (the tokenless helper is removed). LangChain claims and releases in one frame via aiguard_context. Strands, whose own hooks split claim and release, keeps the handle on invocation_state.

Smaller fixes

  • The Anthropic .after event skips streamed requests: a with_raw_response stream reaches it with its body unread, and converting it raised and reported a converter error on every call.
  • is_aiguard_context_active has an early return and a plain loop (it runs on every provider call).
  • Contributor guide (.cursor/rules/ai-guard.mdc) documents which phase each framework claims and each provider check reads, never claiming a phase you do not evaluate, the AI Guard-owned LangChain wrappers, why the stream buffer sits at _stream, and release by handle (never add AI Guard state to a contrib).

Out of scope

Pre-existing issues raised in review are tracked separately in Jira and are not fixed here: REQUEST claimed when LangChain skips its request check, the Strands invocation-wide RESPONSE claim over Strands' own models, the Strands structured_output_async claim leak, the async raw-response parse() coroutine, async evaluation blocking the event loop, Strands Graph parallel nodes sharing one handle, and making the phase a required argument. The double request evaluation when stream() falls back to invoke() (APPSEC-70482) no longer happens here: with the .stream.before listeners removed, that path is evaluated once, by .generate.before.

Testing

  • Context primitive (test_context.py): phase scoping; cross-task release (no raise, clears the claiming task, keeps unrelated claims); a generator closed from another task releases its claim.
  • LangChain (test_langchain.py): streamed chat (sync, async) and LLM responses evaluated before delivery, with the evaluated content checked against the delivered chunks; a response block delivers zero chunks; with_fallbacks does not swallow a block; a model with no provider integration is covered; a streaming AgentExecutor (stream_runnable=True) evaluates its tool call once; streamed tool calls marked evaluated for the agent hooks; claims observed at the provider's HTTP request for streamed and non-streamed calls; the caller's loop body is unclaimed with and without the buffer (APPSEC-70480); a _generate returning with an inner stream still open leaves no claim; an astream() closed from another task leaves no claim; conversion errors fail open.
    New in this revision: callback tokens and on_llm_end arrive only after the verdict, and none arrive on a block; astream_events() on a bare model emits every stream event, after the verdict; a chain's astream_events() emits no model output on a block; invoke(stream=True) is evaluated once, before its tokens (APPSEC-70474); tokens a model reports from inside _stream / _astream wait for the verdict and are dropped on a block (chat sync, async and LLM); a super()._stream subclass is buffered once; one claim per stream, with and without the buffer; a stream under an outer framework claim is still evaluated; the one-call chunk merge matches chunk addition; an inner model streamed by a router model is buffered on its own, so its callbacks see nothing on a block; _stream read directly is buffered for classes defined before and after patching; unpatch removes the hook and the wrappers.
    Review follow-ups: a _stream read directly evaluates its request and blocks before the model is read (deny, abort); a cancelled buffered _astream closes the provider stream; a runtime-defined model class is garbage-collected after buffering; a stream() that falls back to invoke() (no _stream, disable_streaming=True) evaluates its request once (APPSEC-70482); the provider iterator is closed after a drain and after a failure; a subclass that delegates to super()._stream() after its first chunk is evaluated once; native protocol events (_stream_chat_model_events, via stream_events(version="v3")) are buffered, delivered on allow, blocked on deny, and a denied request blocks before the model is read. All of them fail on the previous code.
  • Provider buffer (test_streaming.py): unchanged coverage of evaluate-then-replay and blocking.
  • Anthropic: the after event leaves a streamed raw response untouched and reports no converter error.
  • Strands: cross-task release through the real hooks.

Cassettes: 2 files in tests/cassettes/openai/ for the OpenAI-patched streaming requests (stream_options.include_usage changes the body); response bodies reuse a recorded stream plus a usage chunk.

Local results on the final code:

Suite Result
aiguard::ai_guard_langchain langchain 1.x (py3.13) 125 passed, 7 skipped
aiguard::ai_guard_langchain langchain 0.2.17 (py3.12) 116 passed, 16 skipped
aiguard::ai_guard_langchain langchain 0.1.20 (py3.9) 112 passed, 20 skipped
aiguard::ai_guard_openai 206 passed
aiguard::ai_guard_anthropic 115 passed
aiguard::ai_guard_strands 132 passed
aiguard::ai_guard_api 395 passed
aiguard::ai_guard_litellm_guardrail 70 passed
llmobs::langchain 118 passed, 1 skipped

Skips are the langchain-version-gated agent tests, plus the astream_events(version="v2") tests on langchain-core 0.1 and the native protocol-event tests before langchain-core 1.4. tests/llmobs/test_base_stream_handler.py: 94 passed locally (the .stream.finally pairing test was removed with the event); its async tests need pytest-asyncio, which the local llmobs venv lacks (CI covers them).

Lint clean: fmt, typing, spelling; ast-grep's no-rst-double-backticks warnings in the touched packages dropped from 58 to 49, none on added lines.

Risks

Behaviour change for streamed LangChain responses when DD_AI_GUARD_ANALYZE_STREAM_RESPONSES_ENABLED=true. They are now buffered and evaluated where previously they streamed through unevaluated: higher time-to-first-token, for the caller and for LangChain callbacks (astream_events(), LangGraph messages streaming, models streaming inside invoke()), and one additional evaluation per streamed LangChain call. The flag is off by default, so nobody is affected without opting in. Called out in the release note.

The LangChain contrib change is removals only: it no longer dispatches the AI Guard-only .finally / .stream.started / .stream.finally events (AI Guard was their only listener) and no longer carries AI Guard options. _context.py is private; its no-argument defaults are kept for existing callers and will be removed in a follow-up.

AI Guard wraps LangChain model methods directly, as it already does for agent plan and langgraph ToolNode. The _stream / _astream wrappers and the __init_subclass__ hook are installed at langchain.patch and removed on unpatch. Core ExecutionContexts were considered instead, but a context entered at stream start and exited at finalize would change the caller's current core context for the whole loop, and would stay stuck in the starting task when an async stream is finalized in another task.

Additional Notes

The OpenAI buffered-stream wrappers are installed at openai.patch time only if the flag is already on (existing behaviour); the OpenAI-patched LangChain fixture enables it before patching. The LangChain model-stream buffer checks the flag per call.

🤖 Generated with Claude Code

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Codeowners resolved as

Resolved from the full PR diff against main using the target branch CODEOWNERS file.
CODEOWNERS team requests not listed below are not required by the current file set.

.cursor/rules/ai-guard.mdc                                              @DataDog/asm-python
ddtrace/aiguard/_context.py                                             @DataDog/asm-python
ddtrace/aiguard/_listener.py                                            @DataDog/asm-python
ddtrace/aiguard/_streaming.py                                           @DataDog/asm-python
ddtrace/aiguard/integrations/_anthropic.py                              @DataDog/asm-python
ddtrace/aiguard/integrations/_langchain.py                              @DataDog/asm-python
ddtrace/aiguard/integrations/_openai_chat.py                            @DataDog/asm-python
ddtrace/aiguard/integrations/_openai_responses.py                       @DataDog/asm-python
ddtrace/aiguard/integrations/strands.py                                 @DataDog/asm-python
ddtrace/contrib/internal/langchain/patch.py                             @DataDog/ml-observability
ddtrace/contrib/internal/langchain/utils.py                             @DataDog/ml-observability
docs/configuration.rst                                                  @DataDog/python-guild
releasenotes/notes/aiguard-phase-scoped-collision-avoidance-5a61754799bdbf8e.yaml  @DataDog/apm-python
scripts/check_constant_log_message.py                                   @DataDog/python-guild
tests/aiguard/anthropic/test_anthropic.py                               @DataDog/asm-python
tests/aiguard/langchain/conftest.py                                     @DataDog/asm-python
tests/aiguard/langchain/test_langchain.py                               @DataDog/asm-python
tests/aiguard/openai/conftest.py                                        @DataDog/asm-python
tests/aiguard/openai/test_context.py                                    @DataDog/asm-python
tests/aiguard/openai/test_streaming.py                                  @DataDog/asm-python
tests/aiguard/strands_hooks/test_strands.py                             @DataDog/asm-python
tests/cassettes/openai/openai_chat_completions_post_6a2de228.json       @DataDog/ml-observability
tests/cassettes/openai/openai_chat_completions_post_86aa04cb.json       @DataDog/ml-observability
tests/llmobs/test_base_stream_handler.py                                @DataDog/ml-observability

@cit-pr-commenter-54b7da

Copy link
Copy Markdown

Circular import analysis

⚠️ Existing circular imports

There are 1 circular imports that already exist on the base branch and have not been changed by this PR.

ddtrace.errortracking._handled_exceptions.bytecode_injector -> ddtrace.errortracking._handled_exceptions.callbacks -> ddtrace.errortracking._handled_exceptions.collector -> ddtrace.errortracking._handled_exceptions.bytecode_reporting -> ddtrace.errortracking._handled_exceptions.bytecode_injector

@cit-pr-commenter-54b7da

cit-pr-commenter-54b7da Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Dependency direction analysis

⚠️ Existing dependency direction violations

There are 201 dependency direction violations that already exist on the base branch and have not been changed by this PR.

Show existing violations (showing 5 of 201 highest severity)
ddtrace.internal.tracemethods -×-> ddtrace.trace  (internal-core -> product:tracing, score=132)
ddtrace.llmobs._telemetry -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)
ddtrace.llmobs._integrations.claude_agent_sdk -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)
ddtrace.llmobs._integrations.crewai -×-> ddtrace.trace  (product:llmobs -> product:tracing, score=130)
ddtrace.debugging._signal.model -×-> ddtrace.trace  (product:debugging -> product:tracing, score=130)

To see all violations, download the layers-base.json and layers-pr.json artifacts from this CI job and run:

uv run --script scripts/import-analysis/layers.py compare layers-base.json layers-pr.json

@datadog-datadog-prod-us1-2

datadog-datadog-prod-us1-2 Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Tests

✅ All CI checks and tests passed.

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: 0bf5670 | Docs | View more details | Give us feedback!

@avara1986
avara1986 force-pushed the avara/aiguard-phase-scoped-collision-avoidance branch 2 times, most recently from 8a41031 to 7d263a4 Compare September 21, 2026 13:09
Base automatically changed from avara/aiguard-langchain-response-evaluation to main September 25, 2026 09:24
avara1986 and others added 2 commits September 25, 2026 11:56
…onse phases

AI Guard used one all-or-nothing depth counter to stop a framework and a
provider integration from evaluating the same call twice. While it was
raised, provider listeners skipped both their request and their response
evaluation.

LangChain streaming raised that counter on iteration entry, because it
evaluates the request itself. But it has no stream after-event, so it
could not evaluate the response -- and the counter had already switched
off the provider's buffered-stream evaluation, which was the only thing
left that could. A streamed LangChain response was therefore scanned by
nobody, even with DD_AI_GUARD_ANALYZE_STREAM_RESPONSES_ENABLED=true.

Track the two phases independently so a framework suppresses only what it
actually covers:

  LangChain generate/agenerate  REQUEST + RESPONSE  (evaluates both)
  LangChain streaming           REQUEST only        (provider covers the response)
  Strands                       REQUEST + RESPONSE  (before/after model call)

Provider before-listeners now guard on REQUEST, after-listeners and the
buffered stream on RESPONSE.

set_aiguard_context_active still returns a pairing handle, and no
arguments still claims every phase, so existing callers and the thread /
task isolation and nesting tests keep exercising the new implementation
unchanged.

Also make reset_aiguard_context_active tolerate a token created in a
different Context. Two phases means two tokens and two chances for
ContextVar.reset to raise ValueError into a framework's cleanup path;
it now falls back to a plain decrement, which is correct whether the
context inherited the claim or never saw it. That fixes the Strands
before/after-invocation crash in APPSEC-70282 for every caller.

APPSEC-70286
APPSEC-70282

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The phase-scoping tests only observed the context flags, so nothing proved
a streamed LangChain response is actually evaluated. Run ChatOpenAI.stream
with the OpenAI integration patched underneath and assert the buffered
stream evaluates the response, blocks it before any chunk is delivered, and
stays off when the flag is disabled.

The OpenAI integration injects stream_options.include_usage, so the request
bodies differ from the existing recordings; the two new cassettes reuse a
recorded stream response.

APPSEC-70286

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@avara1986
avara1986 force-pushed the avara/aiguard-phase-scoped-collision-avoidance branch from 7d263a4 to bf3cb54 Compare September 25, 2026 10:33
avara1986 and others added 2 commits September 25, 2026 12:52
The Phase import moved _report_converter_error's add_error_log call from
line 59 to 60, so its line-based exemption in the constant-log checker no
longer matched and the error-log-check CI job failed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@avara1986
avara1986 marked this pull request as ready for review September 25, 2026 13:33
@avara1986
avara1986 requested review from a team as code owners September 25, 2026 13:33
@avara1986
avara1986 requested review from florentinl and removed request for a team September 25, 2026 13:33
@avara1986

Copy link
Copy Markdown
Member Author

@codex review

@avara1986
avara1986 marked this pull request as draft September 25, 2026 13:34

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 5282bf50cd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread ddtrace/aiguard/_context.py Outdated
@avara1986
avara1986 marked this pull request as ready for review September 28, 2026 08:09

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0b1d72be7a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread ddtrace/aiguard/_context.py
Comment thread ddtrace/aiguard/integrations/_langchain.py Outdated
avara1986 and others added 4 commits September 28, 2026 10:59
Carry the .stream.started claim handle on the stream handler so
.stream.finally releases it even when another asyncio task finalizes the
stream. Update the AI Guard guide for phase-scoped claims.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…coped-collision-avoidance

# Conflicts:
#	.cursor/rules/ai-guard.mdc
…ease generate claims by handle

Evaluate streamed LangChain responses in a buffer installed on
BaseChatModel/BaseLLM stream/astream instead of the provider's buffer, so
LangChain's per-chunk read timeout still applies, blocks raise the
BaseException-based AIGuardAbortError that with_fallbacks cannot swallow,
streamed tool calls are recorded as evaluated for the agent hooks, and any
provider is covered. LangChain streams claim both phases again.

Generate now carries its claim handle from .before to .finally in a
per-call state dict; the tokenless reset_aiguard_context_active_current is
removed. The Anthropic after event skips streamed raw responses, the
stream handler releases its claim even if on_span_finish raises, and
is_aiguard_context_active is fast again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@avara1986
avara1986 marked this pull request as draft September 28, 2026 14:11
…e contrib product-agnostic

AI Guard now takes and releases its LangChain claims in one frame of its
own wrappers around BaseChatModel/BaseLLM generate, agenerate, stream and
astream: generate holds the claim for the whole call, streams only while
each chunk is pulled. No claim handle has to travel through the contrib,
so the LangChain contrib drops the AI Guard state dicts, the
aiguard_*_event options and the finally/started events that existed only
to release claims. It keeps its product-agnostic before/after events.

Claiming per chunk also leaves the caller's loop body unclaimed, so
direct SDK calls made inside a LangChain stream loop are evaluated
(APPSEC-70480).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
avara1986 and others added 3 commits September 29, 2026 15:31
… the callbacks

The stream buffer sat on stream() / astream(), above the point where LangChain
hands tokens to callbacks, so callback handlers, astream_events() and LangGraph
messages streaming saw the response before the verdict, and astream_events() on
a bare model emitted no stream events with stream analysis on.

Buffer around each model class's own _stream / _astream instead, installed per
class on first use. Tokens a model reports to the run manager from inside
_stream are deferred until the verdict. This also covers streaming inside
invoke() (APPSEC-70474); .generate.after skips a payload the buffer already
evaluated in the same call. One claim per stream instead of per chunk, and the
chunks are merged in one add_ai_message_chunks call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rs at patch time

The nested-buffer marker was a plain flag, so any other model streamed while a
buffer read its model (a router model's inner model) skipped its own buffer and
its callbacks got tokens before the verdict. The marker now holds the model being
read, and only a super()._stream call on that model skips the buffer.

Buffers were installed on first use of generate or stream, so a path reading
_stream directly (stream_events v3) on a fresh class was not buffered. Install
them at langchain.patch for existing subclasses and from an __init_subclass__
hook for later ones; unpatch removes both. The stream / astream wrappers only
installed buffers and are gone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@avara1986
avara1986 marked this pull request as ready for review September 30, 2026 07:20

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 04cf299cd2

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread ddtrace/aiguard/integrations/_langchain.py
Comment thread ddtrace/aiguard/integrations/_langchain.py Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c9af054706

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread ddtrace/aiguard/integrations/_langchain.py Outdated
avara1986 and others added 4 commits October 1, 2026 10:40
…ffer install lock

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…asses can be collected

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ose cancelled async drains

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f36af78a71

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread docs/configuration.rst
…er request evaluation

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6970afdea9

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread ddtrace/aiguard/integrations/_langchain.py Outdated
Comment thread ddtrace/aiguard/integrations/_langchain.py
…n across reads, buffer native LangChain protocol events

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0bf5670516

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +338 to +340
in_buffer = _BUFFERED_MODEL.set(instance)
try:
item = next(iterator, _END)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Claim phases around every underlying stream read

When stream analysis is disabled and a model yields a prelude before lazily invoking its provider on a later next()—for example, a subclass that delegates to a provider-backed super()._stream() after its first yield—this loop sets only _BUFFERED_MODEL; the phase claim ended after the first read. Provider hooks consult Phase.REQUEST/Phase.RESPONSE, not _BUFFERED_MODEL, so the request is evaluated again and a provider response verdict can be raised after the prelude already reached the caller. The fresh evidence beyond the resolved same-model-suppression thread is that the new per-read guard still omits aiguard_context; wrap every underlying sync and async read in the phase claim while keeping the caller's yield outside it. .cursor/rules/ai-guard.mdcL174-L177

Useful? React with 👍 / 👎.

Comment on lines +189 to +194
for klass in model_cls.__mro__:
if klass is base or klass in _buffered_model_classes or not issubclass(klass, base):
continue
walked.append(klass)
for name, wrapper in _STREAM_METHODS if is_chat else _STREAM_METHODS[:2]:
if name in klass.__dict__:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Wrap stream implementations inherited from non-model mixins

When a concrete model obtains _stream or _astream from a mixin that is not itself a BaseChatModel/BaseLLM subclass—for example, class Model(StreamMixin, BaseChatModel)—Model.__dict__ has no method to wrap and this condition skips the mixin that owns it. The code then marks Model as buffered, so it is never retried, and with stream analysis enabled a model using an otherwise unsupported provider can deliver tokens without any LangChain response verdict. Install a wrapper on the concrete model for its effective inherited stream method, or otherwise record and wrap non-model MRO owners. .cursor/rules/ai-guard.mdcL200-L207

Useful? React with 👍 / 👎.

Comment on lines +549 to +550
if response_messages and _evaluate_langchain_response(client, request_messages, response_messages):
_record_buffer_evaluated(request_messages, response_messages)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Mark native-event tool calls as already evaluated

When _stream_chat_model_events or its async twin produces a tool-call response whose output message is then executed through an instrumented legacy agent or ToolNode, this path records only the payload in _BUFFER_EVALUATED; for a direct stream_events() call that context variable is None, and no fingerprint is attached to the replayed output. Unlike the ordinary chunk path above, the tool execution hook therefore cannot recognize the already-approved call and evaluates it a second time, adding duplicate service requests and potentially raising a later verdict for a call that has already passed response evaluation. Propagate the evaluated tool-call fingerprints through the replayed native events/output message as the chunk path does.

Useful? React with 👍 / 👎.

@avara1986
avara1986 requested a review from Yun-Kim October 2, 2026 11:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants